Papers with natural language processing techniques
BiomedCurator: Data Curation for Biomedical Literature (2022.aacl-demo)
Copied to clipboard
Mohammad Golam Sohrab, Khoa N.A. Duong, Ikeda Masami, Goran Topić, Yayoi Natsume-Kitatani, Masakata Kuroda, Mari Nogami Itoh, Hiroya Takamura
| Challenge: | BiomedCurator uses state-of-the-art natural language processing techniques to extract structured data from scientific articles. |
| Approach: | They propose a web application that extracts structured data from PubMed and ClinicalTrials.gov . the application uses a combination of natural language processing techniques and a pattern-based extraction approach . |
| Outcome: | The proposed system extracts the structured data from PubMed and ClinicalTrials.gov datasets. |
Supporting Complaints Investigation for Nursing and Midwifery Regulatory Agencies (2021.acl-demo)
Copied to clipboard
Piyawat Lertvittayakumjorn, Ivan Petej, Yang Gao, Yamuna Krishnamurthy, Anna Van Der Gaag, Robert Jago, Kostas Stathis
| Challenge: | Fig. 1 illustrates the major components and workflow of our proposed system to improve the efficiency of complaints investigation for nursing and midwifery regulators. |
| Approach: | They propose a decision support system that uses machine learning and natural language processing techniques to process complaints and predict their risk level. |
| Outcome: | The proposed system uses state-of-the-art machine learning and natural language processing techniques to process complaints and predict risk levels. |
Joint Dialogue Topic Segmentation and Categorization: A Case Study on Clinical Spoken Conversations (2023.emnlp-industry)
Copied to clipboard
| Challenge: | Utilizing natural language processing in clinical conversations is effective to improve the efficiency of workflows for medical staff and patients. |
| Approach: | They propose a model for dialogue segmentation and topic categorization that integrates natural language processing techniques into a joint model. |
| Outcome: | The proposed model improves on follow-up calls for diabetes management and reduces computational complexity and cost. |
Generating Image Captions in Arabic using Root-Word Based Recurrent Neural Networks and Deep Neural Networks (N18-4)
Copied to clipboard
| Challenge: | Existing studies on image caption generation in English focus on Western languages, ignoring Semitic and Middle-Eastern languages like Arabic, Hebrew, Urdu and Persian. |
| Approach: | They propose to leverage the critical dependency of Arabic to generate Arabic captions using root-word based Recurrent Neural Network and Deep Neural networks. |
| Outcome: | The proposed model outperforms English-Arabic translated captions on a dataset from newspapers in the Middle East. |
Visualizing Trends of Key Roles in News Articles (D19-3)
Copied to clipboard
| Challenge: | a demonstration system visualizes news trend of key roles based on natural language processing techniques . semantic role labelling and word embeddings can help users understand news topics . |
| Approach: | They propose a system that visualizes the news trend of key roles based on natural language processing techniques. |
| Outcome: | The proposed system analyzes the news trend of key roles using semantic role labelling . it also analyzes how similarities between key roles and news topics change over time . |
DISPUTool 3.0: Fallacy Detection and Repairing in Argumentative Political Debates (2025.acl-demo)
Copied to clipboard
| Challenge: | DISPUTool 3.0 is a web-based application for identifying and fixing fallacious arguments in political debates. |
| Approach: | They propose a web-based application designed to identify and repair fallacious arguments in political debates. |
| Outcome: | The proposed tool is based on the ElecDeb60to20 dataset covering US presidential debates from 1960 to 2020. |
A Natural Approach for Synthetic Short-Form Text Analysis (2024.lrec-main)
Copied to clipboard
| Challenge: | Social media and news sites can be flooded with synthetically generated misinformation via tweets and posts while authentic users can inadvertently spread this text via shares and retweets. |
| Approach: | They propose a method of detecting synthetically generated tweets via a Transformer architecture and incorporate unique style-based features. |
| Outcome: | The proposed method detects synthetically generated tweets using a Transformer architecture and incorporating unique style-based features. |
TROPE: TRaining-Free Object-Part Enhancement for Seamlessly Improving Fine-Grained Zero-Shot Image Captioning (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing approaches to enhance zero-shot abilities in image captioning fail with fine-grained datasets. |
| Approach: | They propose a method to enhance captions with additional object-part details using object detector proposals and natural language processing techniques. |
| Outcome: | The proposed method improves performance on fine-grained datasets and improves on existing methods. |
Enhancing Air Quality Prediction with Social Media and Natural Language Processing (P19-1)
Copied to clipboard
| Challenge: | predicting air quality is a major concern for human health, but the changes of air quality conditions are still difficult to monitor. |
| Approach: | They propose to exploit social media and natural language processing techniques to enhance air quality prediction. |
| Outcome: | The proposed approach improves air quality prediction over baseline that does not use social media by 6.9% to 17.7% in macro-F1 scores. |
ChartThinker: A Contextual Chain-of-Thought Approach to Optimized Chart Summarization (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing methods for chart summarization lack visual-language matching and reasoning ability. |
| Approach: | They propose a method which synthesizes deep analysis based on chains of thought and strategies of context retrieval to improve the logical coherence and accuracy of the generated summaries. |
| Outcome: | The proposed method outperforms 8 state-of-the-art models over 7 evaluation metrics and can significantly reduce time and cognitive resources required. |
Building an Ellipsis-aware Chinese Dependency Treebank for Web Text (L18-1)
Copied to clipboard
| Challenge: | ellipsis is a common linguistic phenomenon that some words are left out as they are understood from the context, especially in oral utterance. |
| Approach: | They propose to use a Chinese dependency treebank to facilitate the parsing of web text . they propose to restore omissions and reserve contexts in the web text to improve dependency parsers . |
| Outcome: | The proposed framework enables the parsing of web text from online microblogs. |
BPM_MT: Enhanced Backchannel Prediction Model using Multi-Task Learning (2021.emnlp-main)
Copied to clipboard
| Challenge: | Backchannel (BC) is a short and quick reaction signal of a listener to a speaker's utterances. |
| Approach: | They propose a model that utilizes lexical information in utterances to enhance backchannel (BC) prediction. |
| Outcome: | The proposed model showed 14.24% performance improvement compared to baseline in the four BC categories: continuer, understanding, empathic response, and No BC. |
Towards Intention Understanding in Suicidal Risk Assessment with Natural Language Processing (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Suicide is a global problem, with one suicide case for every 100 deaths worldwide . social networking sites are an essential forum for communication and information sharing . |
| Approach: | This paper compares natural language processing to suicidal ideation detection and risk assessment . it urges better intention understanding for reliable suicide risk assessment with computational methods . |
| Outcome: | This paper compares the performance of natural language processing to suicidal ideation detection and risk assessment tasks. |
Evaluating Sentence Segmentation in Different Datasets of Neuropsychological Language Tests in Brazilian Portuguese (2020.lrec-1)
Copied to clipboard
| Challenge: | Using automated analysis of connected speech is a promising direction for diagnosing cognitive impairments. |
| Approach: | They propose to use a novel model to segment impaired speech transcriptions . they propose to include a Linear Chain CRF and a self-attention mechanism . |
| Outcome: | The proposed system performs better than the existing model with three new datasets used to diagnose cognitive impairments. |
Towards Processing of the Oral History Interviews and Related Printed Documents (L18-1)
Copied to clipboard
Zbyněk Zajíc, Lucie Skorkovská, Petr Neduchal, Pavel Ircing, Josef V. Psutka, Marek Hrúz, Aleš Pražák, Daniel Soutner, Jan Švec, Lukáš Bureš, Luděk Müller
| Challenge: | a project aims to create an integrated archive of the recordings, scanned documents and photographs from totalitarian regimes in Czechoslovakia . the archive will be accessible online and provide multifaceted search capabilities . |
| Approach: | They propose to use automatic speech recognition and optical character recognition to build an archive of the interviews, scanned documents and photographs. |
| Outcome: | The proposed archive will be accessible online and provide multifaceted search capabilities. |
MIND: A Large-scale Dataset for News Recommendation (2020.acl-main)
Copied to clipboard
Fangzhao Wu, Ying Qiao, Jiun-Hung Chen, Chuhan Wu, Tao Qi, Jianxun Lian, Danyang Liu, Xing Xie, Jianfeng Gao, Winnie Wu, Ming Zhou
| Challenge: | Personalized news recommendation is an important technique for personalized news service. |
| Approach: | They propose to build a large-scale news recommendation dataset from Microsoft News . they demonstrate that news recommendation relies on the quality of news content understanding . |
| Outcome: | The proposed dataset contains 1 million users and more than 160k English news articles, each of which has rich textual content such as title, abstract and body. |
Predicting Anti-Asian Hateful Users on Twitter during COVID-19 (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Xenophobia and polarization have accompanied widespread social media usage in many nations, attracting many researchers. |
| Approach: | They apply natural language processing techniques to characterize Twitter users who began to post anti-Asian hate messages during COVID-19. |
| Outcome: | The results show that it is possible to predict who later posted anti-Asian slurs on Twitter and Reddit. |
Why Swear? Analyzing and Inferring the Intentions of Vulgar Expressions (D18-1)
Copied to clipboard
| Challenge: | Vulgar words are employed in language use for several different functions, including expressing aggression, signaling group identity or the informality of the communication. |
| Approach: | They present a dataset of 7,800 tweets with six categories of vulgarity in which all instances of vulgar words are annotated with one of the six categories. |
| Outcome: | The proposed model can predict the category of a vulgar word based on the immediate context it appears in with 67.4 macro F1 across six classes. |
What to Fuse and How to Fuse: Exploring Emotion and Personality Fusion Strategies for Explainable Mental Disorder Detection (2023.findings-acl)
Copied to clipboard
| Challenge: | Mental health disorders (MHD) are one of the greatest challenges facing our healthcare systems and modern societies in general. |
| Approach: | They integrate and extend the research by conducting extensive experiments with three types of deep learning-based fusion strategies: feature-level fusion, model fusion and task fusion. |
| Outcome: | The proposed techniques show that they can be used to improve mental health detection from textual data. |
CTAP for Italian: Integrating Components for the Analysis of Italian into a Multilingual Linguistic Complexity Analysis Tool (2020.lrec-1)
Copied to clipboard
| Challenge: | Linguistic complexity is a core construct in Second Language Acquisition (SLA) research. |
| Approach: | They present an open source linguistic complexity measurement tool for Italian . they compare it to existing tools for English and germany . |
| Outcome: | The proposed tool is the most comprehensive linguistic complexity measurement tool for italian . it can be used to compare italian texts to multiple other languages in one tool . |